Papers with task-specific models
Efficient Entity Embedding Construction from Type Knowledge for BERT (2022.findings-aacl)
Copied to clipboard
| Challenge: | Existing work has shown advantages of incorporating knowledge graphs (KGs) into BERT for various NLP tasks. |
| Approach: | They propose to integrate knowledge graphs into BERT to train entity embeddings to include rich information of factual knowledge. |
| Outcome: | The proposed models perform very well when combined with context. |
TweetNLP: Cutting-Edge Natural Language Processing for Social Media (2022.emnlp-demos)
Copied to clipboard
Jose Camacho-collados, Kiamehr Rezaee, Talayeh Riahi, Asahi Ushio, Daniel Loureiro, Dimosthenis Antypas, Joanne Boisson, Luis Espinosa Anke, Fangyu Liu, Eugenio Martínez Cámara
| Challenge: | TweetNLP is an integrated platform for natural language processing in social media. |
| Approach: | They propose a Python-based platform for natural language processing in social media that supports a variety of NLP tasks. |
| Outcome: | The proposed platform supports generic focus areas such as sentiment analysis and named entity recognition, as well as social media-specific tasks such as emoji prediction and offensive language identification. |
AdapterHub: A Framework for Adapting Transformers (2020.emnlp-demos)
Copied to clipboard
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, Iryna Gurevych
| Challenge: | AdapterHub framework enables dynamic “stiching-in” of pre-trained adapters for different tasks and languages. |
| Approach: | They propose a framework that allows dynamic "stiching-in" of pre-trained adapters for different tasks and languages. |
| Outcome: | The proposed framework allows dynamic “stiching-in” of pre-trained adapters for different tasks and languages. |
Generation-Distillation for Efficient Natural Language Understanding in Low-Data Settings (D19-61)
Copied to clipboard
| Challenge: | Recent research points to knowledge distillation as a potential solution for NLU tasks. |
| Approach: | They propose a training approach that distills large finetuned LMs into a small network using unlabeled training examples. |
| Outcome: | The proposed approach outperforms BERT training approaches while using 300 times fewer parameters. |
Arcee’s MergeKit: A Toolkit for Merging Large Language Models (2024.emnlp-industry)
Copied to clipboard
Charles Goddard, Shamane Siriwardhana, Malikeh Ehghaghi, Luke Meyers, Vladimir Karpukhin, Brian Benedict, Mark McQuade, Jacob Solawetz
| Challenge: | Open-source language models can merge their parameters to improve performance and versatility without additional training. |
| Approach: | They propose to integrate model checkpoints into powerful multitask models without additional training. |
| Outcome: | the library has facilitated the merging of thousands of models, contributing to some of the world’s most powerful open-source model checkpoints. |
LexSym: Compositionality as Lexical Symmetry (2023.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to generalize compositional models fail to generalise from small datasets. |
| Approach: | They propose a domain-general and model-agnostic formulation of compositionality as a constraint on symmetries of data distributions rather than models. |
| Outcome: | The proposed procedure matches or surpasses state-of-the-art, task-specific models on COGS semantic parsing, SCAN and Alchemy instruction following, and CLEVR-CoGenT visual question answering datasets. |
Speakerly: A Voice-based Writing Assistant for Text Composition (2023.emnlp-industry)
Copied to clipboard
Dhruv Kumar, Vipul Raheja, Alice Kaiser-Schatzlein, Robyn Perry, Apurva Joshi, Justin Hugues-Nuger, Samuel Lou, Navid Chowdhury
| Challenge: | Speakerly TM is a voice-based writing assistance system that works across the different stages of writing. |
| Approach: | They propose a voice-based writing assistance system that helps users with text composition across various use cases such as emails, instant messages, and notes. |
| Outcome: | The proposed system can be used for email, instant messages, and notes. |
Molecular String Representation Preferences in Pretrained LLMs: A Comparative Study in Zero- & Few-Shot Molecular Property Prediction (2025.emnlp-main)
Copied to clipboard
| Challenge: | Molecular property prediction plays a crucial role in medicinal chemistry . traditional machine learning approaches do not involve natural language . |
| Approach: | They compare performance of four state-of-the-art LLMs on molecular property prediction tasks . they find statistically significant zero- and few-shot preferences for InChI and IUPAC names . |
| Outcome: | The proposed model outperforms the current model on molecular property prediction tasks . the model's representation preferences are based on representation granularity, tokenization and prevalence in pretraining corpora . |
UniEDU: Toward Unified and Efficient Large Multimodal Models for Educational Tasks (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing research has focused on plain text, while real-world K-12 scenarios often involve multimodal data. |
| Approach: | They propose a unified language and vision assistant called UniEDU for educational applications . it excels across multiple educational tasks while maintaining strong generalization capabilities . authors propose to use UniEDu for industry-scale deployment . |
| Outcome: | The proposed model excels across multiple educational tasks while maintaining strong generalization capabilities. |
How Many Data Samples is an Additional Instruction Worth? (2023.findings-eacl)
Copied to clipboard
| Challenge: | Recent introduced instruction-paradigm empowers non-expert users to leverage NLP resources by defining a new task in natural language. |
| Approach: | They propose to define a task in natural language without creating task-specific datasets or building models. |
| Outcome: | The proposed model outperforms multitask learning models but is far from state-of-the-art task-specific models. |
Linguistic Knowledge and Transferability of Contextual Representations (N19-1)
Copied to clipboard
| Challenge: | Recent work has explored contextual word representations, which assign each word a vector that is a function of the entire input sequence. |
| Approach: | They compare pretrained word representations with 16 diverse probing tasks to examine their transferability. |
| Outcome: | The pretrained representations are successful across a diverse set of NLP tasks . the models are competitive with state-of-the-art models but fail on fine-grained tasks requiring fine-granular knowledge, the study finds . |
MLLM-I2W: Harnessing Multimodal Large Language Model for Zero-Shot Composed Image Retrieval (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for combining image retrieval are supervised and zero-shot . however, the challenge of mapping pseudo-words to images within the joint image-text embedding space is still a challenge. |
| Approach: | They propose a novel image-text mapping network which converts description-related image information into pseudo-word markers for precise ZS-CIR. |
| Outcome: | The proposed model improves on COCO, CIRR, and Fashion-IQ benchmarks. |
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions (2024.naacl-long)
Copied to clipboard
Siddhant Arora, Hayato Futami, Jee-weon Jung, Yifan Peng, Roshan Sharma, Yosuke Kashiwagi, Emiru Tsunoo, Karen Livescu, Shinji Watanabe
| Challenge: | Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model’s behavior and surpassing performance of task-specific models. |
| Approach: | They adapt a pre-trained automatic speech recognition model to additional tasks using single-token task specifiers. |
| Outcome: | The proposed model can generalize to new datasets and languages for seen task types. |
Task-Agnostic Detector for Insertion-Based Backdoor Attacks (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for textual backdoor detection are task-specific and less effective beyond sentence classification. |
| Approach: | They propose a task-agnostic method for backdoor detection that leverages final layer logits and an efficient pooling technique. |
| Outcome: | TABDet can jointly learn from diverse task-specific models, demonstrating superior detection efficacy over traditional methods. |
Performance-Guided LLM Knowledge Distillation for Efficient Text Classification at Scale (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) face high computational demands at inference time due to high computational costs. |
| Approach: | They propose a cost-effective and high-throughput solution for large language models . PGKD distills the knowledge of LLMs into smaller, task-specific models based on teacher-student knowledge distillation . |
| Outcome: | PGKD outperforms BERT-based models and other knowledge distillation methods on multi-class classification datasets. |
Dynamic Fisher-weighted Model Merging via Bayesian Optimization (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing merging approaches involve scaling the parameters model-wise or integrating parameter importance parameter-wise. |
| Approach: | They propose a method for merging model-based models at the parameter level without training data or joint training. |
| Outcome: | The proposed model merging framework outperforms baseline models on validation sets. |
A Novel Computational Modeling Foundation for Automatic Coherence Assessment (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing models for text coherence assessment rely on a proxy task . however, this approach does not capture the full range of factors contributing to coherency. |
| Approach: | They propose a formal linguistic definition of what makes a discourse coherent and formalize these conditions as respective computational tasks that are jointly trained. |
| Outcome: | The proposed model improves on two human-rated coherence benchmarks. |
Academics Can Contribute to Domain-Specialized Language Models (2024.emnlp-main)
Copied to clipboard
Mark Dredze, Genta Winata, Prabhanjan Kambadur, Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, David Rosenberg, Sebastian Gehrmann
| Challenge: | Commercially available models dominate academic leaderboards, focusing on creating and adapting general-purpose models . however, general- purpose models often underperform in specialized domains, and domain-specific models yield superior results. |
| Approach: | They advocate for a renewed focus on developing and evaluating domain- and task-specific models . they advocate for an adapted or adapted model that can be used to improve academic leaderboard standings . |
| Outcome: | The proposed model can do well on professional and linguistic examinations, college-level knowledge questions, and collections of reasoning tasks. |
Investigating Transfer Learning in Multilingual Pre-trained Language Models through Chinese Natural Language Inference (2021.findings-acl)
Copied to clipboard
| Challenge: | Multilingual transformers have been shown to have remarkable transfer skills in zero-shot settings. |
| Approach: | They investigate cross-lingual transfer abilities of XLM-R for Chinese and English natural language inference using a large scale Chinese dataset. |
| Outcome: | The proposed model trains on Chinese and English natural language inference datasets. |
Structuring Radiology Reports: Challenging LLMs with Lightweight Models (2025.emnlp-main)
Copied to clipboard
Johannes Moll, Louisa Fay, Asfandyar Azhar, Sophie Ostmeier, Sergios Gatidis, Tim C. Lueth, Curtis Langlotz, Jean-Benoit Delbrouck
| Challenge: | Radiology reports lack a standardized format, limiting both interpretability and machine learning applications. |
| Approach: | They propose to use lightweight encoder-decoder models for structuring radiology reports . they compare models with eight open-source LLMs with prompting and in-context learning . |
| Outcome: | The proposed models outperform eight open-source LLMs on a human-annotated test set. |
SynthEval: Hybrid Behavioral Testing of NLP Models with Synthetic Evaluation (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing frameworks for benchmarking in NLP often overestimate performance . however, manually creating a variety of test types requires significant human labor . |
| Approach: | They propose a framework that leverages large language models to generate a wide range of test types . they first generate sentences via LLMs and then identifies challenging examples . |
| Outcome: | The proposed framework overestimates performance on two classification tasks. |
REAR: Reinforced Reasoning Optimization for Event Argument Extraction with Relation-Aware Support (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for EAE restrict integration of relation-level semantics, thereby overlooking the complementary cues from RE. |
| Approach: | They propose a Relation-aware EAE Reinforced optimization framework that integrates relation-level cues from RE into the Large Language Model (LLM) |
| Outcome: | The proposed framework surpasses existing decoder-only methods on the ACE-E, ACE+ and ERE benchmarks. |
Posing Fair Generalization Tasks for Natural Language Inference (D19-1)
Copied to clipboard
| Challenge: | Existing evaluation methods for deep learning semantics rely on naturalistic corpora, but they often fail to support the kind of generalization we are asking for. |
| Approach: | They define and motivate a formal notion of fairness for evaluations of deep learning models for semantics . they then apply it to natural language inference by constructing challenging but provably fair artificial datasets based on the results . |
| Outcome: | The proposed evaluations show that standard neural models fail to generalize in the required ways and even these models do not solve the task perfectly. |
An Investigation of Transfer Learning-Based Sentiment Analysis in Japanese (P19-1)
Copied to clipboard
| Challenge: | Text-based transfer learning techniques can be used to perform downstream tasks. |
| Approach: | They propose to use text-based transfer learning techniques to pre-train a language model in an unsupervised manner and leverage them to perform effective on downstream tasks. |
| Outcome: | The proposed model performs better than task-specific models trained on 3 times as much data and is as effective for language modeling pre-trained on 1/30 of the data. |
Selecting and Merging: Towards Adaptable and Scalable Named Entity Recognition with Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to align large language models with information extraction tasks are costly and not all training data benefits target domains. |
| Approach: | They propose a framework which dynamically Selects and Merges expert models at inference time and combines experts beneficial to target domains. |
| Outcome: | The proposed framework outperforms the unified model by 10% on multiple benchmarks. |
A Thorough Examination of Decoding Methods in the Era of LLMs (2024.emnlp-main)
Copied to clipboard
| Challenge: | Decoding methods are essential for converting language models from next-token predictors into practical task solvers. |
| Approach: | They propose to evaluate decoding methods in general-purpose large language models . they find that decoding method performance is notably task-dependent . |
| Outcome: | The proposed methods perform task-dependently and are influenced by alignment, model size, and quantization. |
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes (2023.findings-acl)
Copied to clipboard
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, Tomas Pfister
| Challenge: | Deploying large language models (LLMs) is difficult because they are memory inefficient and compute-intensive for practical applications. |
| Approach: | They propose a mechanism that fine tunes or distills small models that outperform LLMs . they use human labels to fine tune models or LLM-generated labels to train models . |
| Outcome: | The proposed method outperforms LLMs by using fewer training examples compared to few-shot prompted models using substantially smaller model sizes. |
Data-Efficient Finetuning Using Cross-Task Nearest Neighbors (2023.findings-acl)
Copied to clipboard
| Challenge: | Prior work shows training models on multitask data augmented with task descriptions transfers knowledge to new tasks. |
| Approach: | They propose to use unlabeled target-task data to train models on task descriptions . they use only 2% of the data from the P3 pool without labeled target task data . |
| Outcome: | The proposed model outperforms baseline models on 12 out of 14 datasets . it also provides better initialization than single model on target-task data . |
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning (2024.findings-acl)
Copied to clipboard
| Challenge: | Charts provide visual representations of data and are used for analyzing information, addressing queries, and conveying insights to others. |
| Approach: | They propose a chart-specific vision-language Instruction-following dataset with 191K instructions and a pipeline model that extracts chart data tables and inputs them into a LLM. |
| Outcome: | The proposed model can solve a wide range of chart-related tasks, achieving state-of-the-art results on four tasks. |
Fisher Mask Nodes for Language Model Merging (2024.lrec-main)
Copied to clipboard
| Challenge: | Pre-trained models are ubiquitous in natural language processing, but individual fine-tuned models require significant overhead in multi-task scenarios. |
| Approach: | They propose a method for fine-tuning pre-trained models for Transformers using Fisher information. |
| Outcome: | The proposed method outperforms Fisher-weighted averaging in a fraction of the computational cost. |
Sparsity Makes Sense: Word Sense Disambiguation Using Sparse Contextualized Word Representations (2020.emnlp-main)
Copied to clipboard
| Challenge: | Using sparse word embeddings is highly applicable for word sense disambiguation (WSD) . |
| Approach: | They propose an overcomplete set of semantic basis vectors that allows for sparse word representations. |
| Outcome: | The proposed framework achieves an aggregated F score of 78.8 over five standard word sense disambiguating benchmark datasets. |
Seq2seq is All You Need for Coreference Resolution (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on coreference resolution suggests task-specific models are necessary . a recent line of work that take an alternative approach leveraging advances in seq2seq-based models is needed . |
| Approach: | They propose a pretrained seq2seq transformer to map an input document to a tagged sequence encoding the coreference annotation. |
| Outcome: | The proposed model outperforms or matches the best coreference systems on an array of datasets. |
QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent research has focused on smaller, task-specific models enhanced by distilling knowledge from LLMs, but the diversity and quality of negative knowledge remains understudied. |
| Approach: | They propose a quality-guided contrastive rationale distillation framework that aims to enhance reasoning capabilities through contrastive knowledge learning. |
| Outcome: | The proposed method consistently outperforms existing distillation techniques yielding higher-quality rationales. |
LLMaAA: Making Large Language Models as Active Annotators (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing supervised learning methods in natural language processing require large amounts of data. |
| Approach: | They propose an active learning loop that takes LLMs as annotators and puts them into an active loop to determine what to annotate efficiently. |
| Outcome: | The proposed model outperforms existing models with few-shot performance in two NLP tasks. |
Plug-and-Play Document Modules for Pre-trained Models (2023.acl-long)
Copied to clipboard
Chaojun Xiao, Zhengyan Zhang, Xu Han, Chi-Min Chan, Yankai Lin, Zhiyuan Liu, Xiangyang Li, Zhonghua Li, Zhao Cao, Maosong Sun
| Challenge: | Large-scale pre-trained models have been widely adopted for document-oriented NLP tasks, such as question answering. |
| Approach: | They propose to decouple document encoding from downstream tasks by introducing a document plugin into the backbone of a PTM. |
| Outcome: | The proposed model can encode documents once and for all across different scenarios. |
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning large language models for specialized tasks are costly and time-consuming. |
| Approach: | They propose a framework that locates task-specific neurons via gradient-based attribution and dynamically Elects critical neurons through multi-model importance fusion. |
| Outcome: | The proposed framework reduces harmful response rates while preserving 95% of utility performance. |
Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model Merging (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning large language models fail due to performance degradation . existing methods fail for models fine- tuned with low-rank adaptation . |
| Approach: | They propose to constrain the LoRA subspace prior to fine-tuning to ensure that updates relevant to one task do not adversely shift outputs for others. |
| Outcome: | The proposed method can integrate with most existing merging algorithms, reducing unintended interference among tasks. |
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages (2026.acl-long)
Copied to clipboard
| Challenge: | Existing research has largely overlooked lower-resource languages for automated fact-checking. |
| Approach: | They propose a multilingual CND corpus spanning 18 languages across three resource levels and a small decoder-based language model for CND. |
| Outcome: | The proposed model outperforms prompted LLMs in cross-lingual CND across languages. |
SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for sign language processing have relied on task-specific models, limiting the potential for transfer learning across tasks. |
| Approach: | They propose a self-supervised contextual representation model that adapts masked token prediction objectives to multi-stream visual sign language input. |
| Outcome: | The proposed model adapts masked token prediction objectives to multi-stream visual sign language input, learning to predict multiple targets corresponding to clustered hand, face, and body pose streams. |